Controlling hallucinations at word level in data-to-text generation

نویسندگان

چکیده

Abstract Data-to-Text Generation (DTG) is a subfield of Natural Language aiming at transcribing structured data in natural language descriptions. The field has been recently boosted by the use neural-based generators which exhibit on one side great syntactic skills without need hand-crafted pipelines; other side, quality generated text reflects training data, realistic settings only offer imperfectly aligned structure-text pairs. Consequently, state-of-art neural models include misleading statements –usually called hallucinations—in their outputs. control this phenomenon today major challenge for DTG, and problem addressed paper. Previous work deal with issue instance level: using an alignment score each table-reference pair. In contrast, we propose finer-grained approach, arguing that hallucinations should rather be treated word level. Specifically, Multi-Branch Decoder able to leverage word-level labels learn relevant parts instance. These are obtained following simple efficient scoring procedure based co-occurrence analysis dependency parsing. Extensive evaluations, via automated metrics human judgment standard WikiBio benchmark, show accuracy our effectiveness proposed Decoder. Our model reduce hallucinations, while keeping fluency coherence texts. Further experiments degraded version ToTTo could successfully used very noisy settings.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Controlling Logical Scope in Text Generation

Many applications of Natural Language Generation employ object-oriented knowledge representation formalisms such as loom. Each entity in the domain is represented by an object belonging to a speciied class and possessing attributes suitable for that class. From a logical viewpoint, domain entities correspond to existentially quantiied variables, but since loom and kindred formalisms provide no ...

متن کامل

Controlling Lexical Substitution in Computer Text Generation

Th=s report describes Paul, a computer text generation system desig~ed LO create cohesive text through the use o| lexlcal substitutions. Specihcally, Ihas system is designed Io determmistically choose between provluminahzat0on, superordinate suhstntut0on, and dehmte noun phrase reiterabon. The system identities a strength el antecedence recovery for each of the lex~cal subshtutions, and matches...

متن کامل

Analysing Data-To-Text Generation Benchmarks

Recently, several data-sets associating data to text have been created to train data-to-text surface realisers. It is unclear however to what extent the surface realisation task exercised by these data-sets is linguistically challenging. Do these data-sets provide enough variety to encourage the development of generic, high-quality data-to-text surface realisers ? In this paper, we argue that t...

متن کامل

Syntax and Data-to-Text Generation

With the development of the web of data, recent statistical, data-to-text generation approaches have focused on mapping data (e.g., database records or knowledge-base (KB) triples) to natural language. In contrast to previous grammar-based approaches, this more recent work systematically eschews syntax and learns a direct mapping between meaning representations and natural language. By contrast...

متن کامل

The Secret's in the Word Order: Text-to-Text Generation for Linguistic Steganography

Linguistic steganography is a form of covert communication using natural language to conceal the existence of the hidden message, which is usually achieved by systematically making changes to a cover text. This paper proposes a linguistic steganography method using word ordering as the linguistic transformation. We show that the word ordering technique can be used in conjunction with existing t...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: Data Mining and Knowledge Discovery

سال: 2021

ISSN: ['1573-756X', '1384-5810']

DOI: https://doi.org/10.1007/s10618-021-00801-4